Skip to content

ROCm: compress the embedded GPU code; upgrade llama.cpp b11247 → b11256 - #470

Merged
bernardladenthin merged 3 commits into
mainfrom
ccr-d629238e-hohc22
Sep 29, 2026
Merged

bernardladenthin merged 3 commits into
mainfrom
ccr-d629238e-hohc22

Conversation

@bernardladenthin

Copy link
Copy Markdown
Owner

Summary

  • ROCm GPU code compressed (--offload-compress). The Windows ROCm jllama.dll in llama-*-all-windows-x86-64-jar-with-dependencies.jar was ~1 GB. Each HIP translation unit embeds one code object per GPU target (the whole kernel set, once for each of the 23 architectures), and clang stores those bundles uncompressed by default.
    • The jar hid the size (234 MB zipped), but LlamaLoader extracts the library to the temp dir on every start. In the all-backends fat jar, ROCm is tried right after CUDA, so this happens on nearly every machine without an NVIDIA card.
    • ggml-hip is now compiled with --offload-compress, scoped to that target and its source language (HIP on Linux, CXX on Windows). The HIP runtime decompresses each bundle when the module loads.
    • Upstream llama.cpp does not do this.
  • CI guard: both ROCm jobs run the new .github/verify-hip-offload-compressed.py after the build.
    • It prints the library size and the compressed/uncompressed bundle counts, also into the job summary.
    • It fails on any uncompressed bundle (__CLANG_OFFLOAD_BUNDLE__), or if there is no compressed one at all, so losing the flag reds the job instead of shipping the 1 GB library again.
  • llama.cpp b11247 → b11256: 9 commits, 22 KiB diff, one step. No project source change is needed.
    • #29632 moves several upstream tools to llama_backend_init(), which jllama, TTS and the trainer already call.
    • fs_write_atomic() (#29642) is unused here; the scheduler graph_inputs pass (#29634) and faster GGUF duplicate checks (#29598) are internal.
    • server-schema.cpp and server-task.cpp are untouched, so the request/response contract is unchanged.
    • All nine patches apply unchanged, and the drop-checks show all are still needed.
    • No upstream release.yml change, so the CUDA/ROCm/OpenVINO pins stay.
    • History rows are appended to docs/history/llama-cpp-breaking-changes.md.

Test plan

  • Fresh configure on b11256: all nine patches applied (stamp checked); CPU build of jllama + jllama_test succeeds
  • ctest: 590/590 C++ tests pass
  • NativeLibraryLoadSmokeTest 4/4 (linked build-info matches the b11256 pin)
  • The check script was tested on synthetic files: pass, uncompressed bundle, no bundle, no library, and matches straddling chunk boundaries
  • Not verifiable here (no ROCm toolchain in this environment): the ROCm build with --offload-compress and the resulting DLL/.so size. The two ROCm jobs' new "Verify the GPU code is compressed" step and its job-summary table show both.
  • CI is green on this branch

Related issues / PRs

Refs #468 (b11237 → b11247)

Checklist

  • My commits follow the repository's commit style
  • No security-sensitive changes

🤖 Generated with Claude Code

https://claude.ai/code/session_01FtfgoazykGcQSCR3TYBmTQ


Generated by Claude Code

Nine upstream commits, one reviewable step (22 KiB diff). No project
source change: #29632 moves several tools/examples to llama_backend_init(),
which jllama, TTS and the trainer already call; fs_write_atomic() (#29642)
is unused here; the scheduler graph_inputs pass (#29634) and the faster
GGUF duplicate checks (#29598) are internal. server-schema.cpp and
server-task.cpp are untouched, so the request/response contract is
unchanged. All nine patches apply unchanged and are all still needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FtfgoazykGcQSCR3TYBmTQ
The Windows ROCm jllama.dll was ~1 GB: every HIP translation unit embeds
one code object per GPU target, i.e. the whole kernel set once per
architecture, and clang stores those bundles uncompressed by default.
The jar hid it (234 MB zipped), but LlamaLoader extracts the library to
the temp dir on every start, and in the all-backends fat jar ROCm is
tried right after CUDA - on nearly every machine without an NVIDIA card.

ggml-hip is now compiled with --offload-compress, so each bundle is
stored zstd-compressed (CCOB) and inflated by the HIP runtime at module
load. The option is scoped to the ggml-hip target and to its source
language (HIP on Linux, CXX on Windows, where upstream compiles HIP as
C++).

Both ROCm jobs run .github/verify-hip-offload-compressed.py after the
build: it prints the library size and bundle counts (also into the job
summary) and fails on any uncompressed bundle, or on none compressed,
so a toolchain or upstream change cannot quietly bring the 1 GB library
back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FtfgoazykGcQSCR3TYBmTQ
CLAUDE.md said our target lists add gfx900/gfx906/gfx90c/gfx1153 to
upstream's. That holds for Linux only: upstream's windows-rocm list
already carries gfx1153 (checked at b11247 and b11256), so the Windows
extras are just gfx900/gfx906/gfx90c. Same correction in the Windows
ROCm job's comment. No target list changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FtfgoazykGcQSCR3TYBmTQ
@sonarqubecloud

Copy link
Copy Markdown

@bernardladenthin
bernardladenthin merged commit bb1e9f0 into main Sep 29, 2026
12 of 17 checks passed
@bernardladenthin
bernardladenthin deleted the ccr-d629238e-hohc22 branch September 29, 2026 16:09

This branch had an error being deployed

1 failed deployment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants